Can Audio Large Models Remember Who Is Speaking? New Benchmark VoxMem from Monash University Exposes: All Fail Under 32K Context
The University of Melbourne and the University of New South Wales joint team have introduced a speech AI evaluation benchmark called VoxMem to test whether it can remember the speaker's identity, tone, and background sounds during long hours of conversation, dozens of real dialogues. The team pointed out that previous benchmarks only focus on transcribing text, losing identity, emotion, and environmental cues in the sound, making it impossible to measure speech memory capabilities. VoxMem was designed to fill this gap.